Papers with human expert validation

4 papers
SwiLTra-Bench: The Swiss Legal Translation Benchmark (2025.acl-long)

Copied to clipboard

Challenge: In Switzerland legal translation relies on legal experts who must be both legal experts and skilled translators—creating bottlenecks and impacting effective access to justice.
Approach: They propose a multilingual benchmarking system that evaluates Swiss legal translation systems based on 180K aligned Swiss legal translator pairs . they show frontier models achieve superior translation performance across all document types while specialized translation systems excel specifically in laws but under-perform in headnotes.
Outcome: The proposed model outperforms specialized models in laws but underperform in headnotes.
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly applied to creative domains, yet performance in classical Chinese poetry generation and evaluation remains poorly understood.
Approach: They propose a framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation to evaluate large language models.
Outcome: The proposed framework evaluates state-of-the-art LLMs across multiple dimensions of poetic quality in Tang poetry generation.
Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration (2026.findings-acl)

Copied to clipboard

Challenge: Structured Query Language (SQL) is the cornerstone for data-driven decision-making.
Approach: They propose a benchmark to rigorously evaluate Large Language Models within a dynamic interaction framework.
Outcome: The proposed benchmark aims to rigorously evaluate LLMs within a dynamic interaction framework.
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation (2026.acl-long)

Copied to clipboard

Challenge: Traditional metrics like BLEU and BERTScore fail to capture semantic fidelity in generative text-to-text tasks.
Approach: They propose a cross-examination framework that generates verifiable questions from each text and performs a Cross-exam to derive three interpretable scores: Coverage, Conformity, and Consistency.
Outcome: The proposed framework detects critical errors across translation, summarization and clinical note-generation and human expert validation shows it is reliable without gold references.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations